Papers with image-text matching

8 papers
DEMO: A Statistical Perspective for Efficient Image-Text Matching (2024.naacl-long)

Copied to clipboard

Challenge: Image-text matching is a problem that seeks to connect vision and language through semantic understanding.
Approach: They propose a deep unsupervised hashing-based approach for image-text matching . they characterize each image using multiple augmented views, which are considered as samples .
Outcome: The proposed approach achieves superior performance on image-text matching datasets compared with state-of-the-art methods.
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding (2023.findings-acl)

Copied to clipboard

Challenge: supervised methods for vision-language tasks have been well-studied, but they lack the fine-grained information needed for semantics understanding.
Approach: They propose a framework to take advantage of fine-grained information for zero-shot vision-language learning, covering multiple tasks such as VQA, SNLI-VE, and VCR.
Outcome: The proposed framework outperforms previous zero-shot methods on VQA and achieves substantial improvement on SNLI-VE and VCR.
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Recent work reveals that vision and language models struggle to comprehend fine grained distinctions in images.
Approach: They propose a dataset to assess multimodal models' ability to match objects with their colors.
Outcome: The proposed model performs well in visual questionanswering, text-to-image generation and word-order understanding tasks.
MedICaT: A Dataset of Medical Images, Captions, and Textual References (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing largescale datasets explicitly exclude compound figures . existing systems lack this ability to identify relevant subfigures .
Approach: They propose a dataset of medical images in context that allows figure-to-text alignment . they use captions, inline references and manually annotated subfigures for compound figures .
Outcome: The proposed dataset demonstrates the utility of inline references in image-text matching.
MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets (2024.naacl-long)

Copied to clipboard

Challenge: Existing approaches to augment textual dialogues with retrieved images pose privacy, diversity, and quality constraints.
Approach: They propose a framework to augment text-only dialogues with diverse and high-quality images by using a diffusion model and a feedback loop.
Outcome: The proposed framework is comparable to or better than baselines, with significant improvements in human evaluation, especially against retrieval baselines where the image database is small.
MURAL: Multimodal, Multitask Representations Across Languages (2021.findings-emnlp)

Copied to clipboard

Challenge: Image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages.
Approach: They propose a dual encoder that integrates image-text matching and translation pairs to solve two tasks by learning from billions of pairs.
Outcome: The proposed encoder outperforms ALIGN's cross-modal retrieval performance on well-resourced languages and significantly improves on under-resource languages.
Learning Multimodal Contrast with Cross-modal Memory and Reinforced Contrast Recognition (2024.findings-acl)

Copied to clipboard

Challenge: Using a memory module, we learn multimodal contrast using encoding-decoding paradigm . multimodal information are used in many applications, including news feeding, social media, etc.
Approach: They propose an LLM-based approach for learning multimodal contrast following the encoding-decoding paradigm . they use a memory module with reinforced contrast recognition to enhance learning .
Outcome: The proposed approach outperforms baseline and state-of-the-art studies on four English and Chinese benchmark datasets.
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts (2024.emnlp-main)

Copied to clipboard

Challenge: Multimodal models focus on the correspondence between images and text, but this only covers a subset of real-world interactions.
Approach: They propose an approach to enhance multimodal models by training separate expert models for each type of interaction, such as redundancy present in both modalities, uniqueness in one modality, or synergy that emerges when both . modality is used to capture overlaps in semantic content between images and text, making a strong multi-view redundancies assumption.
Outcome: The proposed approach improves on a sarcasm detection and humor detection task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations